Papers with zero-shot TTS models

2 papers
Zero-Shot Text-to-Speech for Vietnamese (2025.acl-short)

Copied to clipboard

Challenge: Text-to-speech (TTS) synthesis has seen significant advancements in recent years.
Approach: They propose to use PhoAudiobook to curated 941 hours of high-quality audio for Vietnamese text-to-speech models.
Outcome: The proposed model improves on VALL-E, VoiceCraft, and XTTS-V2 models, highlighting their robustness in handling diverse linguistic contexts.
Cross-Domain Audio Deepfake Detection: Dataset and Analysis (2024.emnlp-main)

Copied to clipboard

Challenge: Existing audio deepfake detection datasets are outdated and lack generalization capabilities.
Approach: They construct a new cross-domain audio deepfake detection dataset comprising over 300 hours of speech data that is generated by five advanced zero-shot TTS models.
Outcome: The proposed models achieve 4.1% and 6.5% error rates in the cross-domain ADD dataset generated by five advanced zero-shot TTS models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations